Goto

Collaborating Authors

 visual goal-step inference


Visual Goal-Step Inference using wikiHow

arXiv.org Artificial Intelligence

Procedural events can often be thought of as a high level goal composed of a sequence of steps. Inferring the sub-sequence of steps of a goal can help artificial intelligence systems reason about human activities. Past work in NLP has examined the task of goal-step inference for text. We introduce the visual analogue. We propose the Visual Goal-Step Inference (VGSI) task where a model is given a textual goal and must choose a plausible step towards that goal from among four candidate images. Our task is challenging for state-of-the-art muitimodal models. We introduce a novel dataset harvested from wikiHow that consists of 772,294 images representing human actions. We show that the knowledge Figure 1: An example Visual Goal-Step Inference learned from our data can effectively transfer Task: given a text goal (bake white fish), select the to other datasets like HowTo100M, increasing image (C) that represents a step towards that goal. the multiple-choice accuracy by 15% to 20%.